Papers with Paraphrasing Text Embedding Benchmark
PTEB: Towards Robust Text Embedding Evaluation via Stochastic Paraphrasing at Evaluation Time with LLMs (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing evaluations of sentence embedding models rely on static tests like the Massive Text Embedding Benchmark (MTEB) repeated tuning on a fixed suite can inflate reported performance and obscure real-world robustness. |
| Approach: | They propose a dynamic protocol that generates meaning-preserving paraphrases at evaluation time and aggregates results across multiple runs. |
| Outcome: | The proposed protocol generates meaning-preserving paraphrases at evaluation time and aggregates results across multiple runs. |